Back

BMC Bioinformatics

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match BMC Bioinformatics's content profile, based on 457 papers previously published here. The average preprint has a 0.29% match score for this journal, so anything above that is already an above-average fit.

1
Using Shared Features Improves Metabolite Effect Estimation

Dubey, H. V.; Farage, G.; Sen, S.

2026-08-06 bioinformatics 10.64898/2026.07.31.742124 medRxiv
Top 0.1%
40.0%
Show abstract

External biological knowledge provides valuable information about relationships among metabolites, yet this information is usually not incorporated directly into statistical estimation procedures. Most existing approaches estimate metabolite effects independently, ignoring known biochemical structure such as shared subclasses and pathway membership. We propose a Bayesian hierarchical framework that improves metabolite effect estimates by incorporating external biological information describing relationships among metabolites. The proposed method improves metabolite-specific estimates by allowing related metabolites to borrow information from one another while preserving metabolite-level inference. We evaluate the methodology using simulation studies across a range of sample sizes and heterogeneity regimes together with three metabolomics applications involving distinct biological annotation structures. Across both simulated and real datasets, incorporating external biological information consistently improves metabolite effect estimation. Gains are most pronounced when sample sizes are small and metabolite classes are informative, i.e. more homogenous within classes.

2
LightAlign: a lightweight pairwise aligner for memory-constrained HiFi read assembly

Liu, J.; Zhang, J.

2026-08-05 bioinformatics 10.64898/2026.07.30.741934 medRxiv
Top 0.1%
38.9%
Show abstract

IntroductionCurrent de novo genome assembly tools often demand substantial memory resources, and their execution typically relies on high-performance computing (HPC) clusters. This dependency limits their use in resource-constrained settings. Furthermore, mainstream third-generation sequencing assembly and alignment tools usually require explicit detection of overlap regions between reads, a process that often entails significant computational and storage overhead. ResultsTo address this issue, we developed LightAlign, a lightweight alignment tool for HiFi data that innovatively uses sequence-derived fuzzy features and reduces the peak memory usage during overlap detection. ConclusionsWhen combined with miniasm, LightAlign generated bacterial draft assemblies while maintaining peak memory usage below 1 GB and completed overlap generation for the tested eukaryotic datasets within 1.88 GB RAM.

3
User-friendly transcriptomic data analysis with ArrayAnalysis

Koetsier, J.; Cinar, O.; Willighagen, E. L.; Ammar, A.; Karthik, V.; Jennen, D.; Evelo, C. T.; Curfs, L. M. G.; Reutelingsperger, C. P.; Bahram Sangani, N.; Eijssen, L. M. T.

2026-07-18 bioinformatics 10.64898/2026.07.13.738193 medRxiv
Top 0.1%
26.2%
Show abstract

Transcriptomic profiling has become a cornerstone of modern biomedical research. To make transcriptomic analyses accessible to a broader scientific community, specifically including researchers with limited bioinformatics expertise, we introduced ArrayAnalysis in 2013 as a user-friendly web-based application for microarray data analysis. We now present a major update (https://arrayanalysis.org), introducing a strongly interactive platform that facilitates the dedicated exploration and analysis of both microarray and RNA-seq data, and allows for the generation of publication-ready outputs. Users can perform key analysis steps, including data pre-processing and quality control, differential expression analysis, and gene set analysis, via a sequential, interactive workflow. At each step, the application provides interactive visualizations accompanied by information pages to support interpretation. Users can dynamically adjust figure layouts and colour palettes and export figures as vector graphics and high-resolution raster images. For non-expert users, ArrayAnalysis offers step-by-step guidance to support correct usage and facilitate learning, while for experienced bioinformaticians, it provides a streamlined and flexible workflow ideal for large-scale analyses requiring efficient and consistent processing. ArrayAnalysis is available both as a web application and for local deployment as a desktop application, Docker image, or R package, making it suitable for diverse computational environments, user groups, and analytical purposes. Together, ArrayAnalysis empowers a broad community of biomedical researchers to unlock the full potential of transcriptomic data. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=96 SRC="FIGDIR/small/738193v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@1072123org.highwire.dtl.DTLVardef@11095e0org.highwire.dtl.DTLVardef@1dfaee7org.highwire.dtl.DTLVardef@53d31e_HPS_FORMAT_FIGEXP M_FIG C_FIG

4
GCM: metric-guided clustering by genetic algorithm for correlation-defined modules

Madrigal-Roca, L. J.; Kelly, J. K.

2026-07-20 bioinformatics 10.64898/2026.07.15.738632 medRxiv
Top 0.1%
23.2%
Show abstract

Gene co-expression analyses identify "good" modules by a correlation criterion. However, standard pipelines detect modules with greedy algorithms that optimize other quantities and only measure correlation afterwards. We present a method called Genetic Clustering by Metric(GCM hereafter), an open-source Python tool that closes this gap by treating module detection as maximum-likelihood inference and solving it globally. In GCM, the correlation objective is, up to a constant and the sample-size factor, the profile log-likelihood of an explicit generative model: a block-diagonal one-factor Gaussian in which each module is a single regulator with equal-magnitude {+/-} loadings. This bases model selection on a principled footing through a genuine BIC/AIC in correlation space. GCM maximizes this likelihood with a memetic genetic algorithm: a population-based search hybridized with a greedy local refinement that reassigns genes after the fact, a move the agglomerative clustering at the core of co-expression pipelines cannot make. Across a replicated noise sweep, GCM reproducibly surpasses hierarchical correlation clustering and k-means with the lowest variance, and an ablation shows the local-search step is responsible; the advantage persists when the number of modules is unknown and when unstructured genes must be ignored. GCM faithfully optimizes geometric indices on the Iris benchmark dataset. For a breast-cancer RNA-seq it recovers coherent modules that predict tumor-versus-normal status. GCM depends only on NumPy and SciPy and exposes one swappable-metric interface with single- and multi-objective modes. Author summaryWhen biologists group genes by how similarly they are expressed, they usually run a standard clustering method and then score the result with a separate quality measure. The method, however, was never trying to do well on that measure because it optimizes its own internal objective. We built a tool, GCM, that removes this gap: the user picks the quality measure they actually care about, and the tool searches directly for the grouping that scores best on it. The search is performed by a genetic algorithm, a population-based optimizer that mixes and mutates candidate groupings over many generations. GCM includes a purpose-built score for "modules" of co-expressed genes, as well as several widely used geometric scores, and it can balance two competing scores at once to choose how many groups the data support. We show on synthetic data with a known answer, on a textbook dataset, and on real expression data that the tool recovers the intended structure and lets researchers make explicit, and optimize for, their own definition of a good cluster.

5
nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality

Wyatt, C. D. R.; Duarte Frutos, F.; Turner, S. D.; Cerqueira De Araujo, A.; Rashid, U.; Begley, V.; Sumner, S.

2026-08-20 bioinformatics 10.64898/2026.08.20.745971 medRxiv
Top 0.1%
22.0%
Show abstract

The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.

6
Detecting CYP2C19 deletions from genotyping array signals using neural networks

Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.

2026-08-25 bioinformatics 10.64898/2026.08.21.746170 medRxiv
Top 0.1%
18.9%
Show abstract

Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.

7
Interpretable biomarker discovery from small-sample microarray datasets using XGBoost rank aggregation and SVM-RFECV

Prapty, M. M.; Rahman, M. S.

2026-08-21 bioinformatics 10.64898/2026.08.13.744652 medRxiv
Top 0.1%
18.8%
Show abstract

MotivationHigh-dimensional microarray datasets remain valuable for cancer biomarker discovery, but their small sample sizes make robust and interpretable feature selection challenging. Efficient workflows are needed to derive compact gene signatures while preserving biological interpretability. ResultsWe developed a two-stage biomarker-discovery workflow that combines cross-validated XGBoost rank aggregation with support vector machine recursive feature elimination and cross-validation (SVM-RFECV) to identify compact candidate biomarker panels. The workflow was evaluated on 21 public binary and multiclass microarray datasets using repeated stratified cross-validation for internal validation. Across the dataset collection, the selected panels demonstrated strong internal discriminative performance while remaining sufficiently compact for downstream biological interpretation. SHAP analysis identified dataset- and class-specific discriminative genes, and functional enrichment analysis supported the biological coherence of representative consensus signatures. The proposed workflow provides an interpretable and reproducible framework for candidate biomarker discovery from small-sample microarray datasets. AvailabilitySource code and processed outputs are freely available at https://github.com/mashiyat-mahjabin-prapty/microarray-feature-selection.

8
Hobrac: a reference-guided workflow for genome comparison and synteny visualization

Istace, B.; Denoeud, F.; Teodori, E.; Chorba, N.; Aury, J.-M.

2026-07-22 bioinformatics 10.64898/2026.07.17.739168 medRxiv
Top 0.2%
18.7%
Show abstract

Whole-genome comparison is fundamental for validating genome assemblies and investigating genome evolution, yet identifying suitable reference genomes and interpreting chromosome-scale synteny from often noisy nucleotide alignments remain challenging. We introduce Hobrac, an automated workflow that addresses these two major bottlenecks by combining automated reference genome selection with gene-based structural comparisons. Starting from a genome assembly and its taxon identifier, Hobrac identifies suitable reference genomes, complements nucleotide alignments with conserved BUSCO orthologues, and generates publication-quality visualizations. The workflow produces dotplots, ribbon-plots and synteny visualization that can be explored interactively or offline. Hobrac is freely available at https://github.com/Genoscope-LBGB/hobrac.

9
When the Background Matters: Topic-Dependent reference lists in GWAS and Exome Analyses

Timoney, B.; Guasoni, P.; Zade, K.; Bach, S.; Tropea, D.

2026-08-21 bioinformatics 10.64898/2026.08.14.744838 medRxiv
Top 0.2%
18.3%
Show abstract

Gene Ontology (GO) Biological Process overrepresentation analysis is widely used to interpret gene lists from genetic studies, yet results depend critically on the background (universe/reference list) against which enrichment is tested. This paper examines how genome-exome background mismatch alters GO Biological Process significance and induces annotation-driven bias. First, Monte Carlo simulations across multiple input gene list sizes show that enrichment p-values shift systematically when lists sampled from an exome-like universe are tested against a genome background (and vice versa), producing both inflation and deflation of significance depending on GO term composition; these shifts increase with gene list size. Second, applied analyses of gene lists derived from Genome-Wide Association Studies (GWAS) and Whole Exome Studies (WES) across brain, immune, and metabolic domains demonstrate that background choice changes the set of significant GO IDs, yielding reference-specific terms consistent with both Type I errors (false positives) and Type II errors (false negatives). Because genome backgrounds are commonly used by default, the practical risk is greatest when WES-derived lists are analyzed with genome reference lists. To support reproducible best practice, we provide a simple command set for selecting and documenting study-appropriate backgrounds and for assessing sensitivity of GO Biological Process results to the chosen universe.

10
SyntenyPair Explorer: an installation-free, browser-based tool for interactive pairwise genome synteny visualization

Gibbons, J. G.

2026-07-27 bioinformatics 10.64898/2026.07.23.740353 medRxiv
Top 0.2%
17.6%
Show abstract

Comparisons of genome structure between related organisms are central to understanding genome evolution, gene family dynamics, and the genomic basis of phenotypic variation. Synteny, the conserved co-localization of genes along chromosomes, is most readily interpreted visually, yet many existing synteny visualization tools require local software installation, command-line proficiency, and/or dedicated server infrastructure, and produce static images that cannot be explored interactively. Here, I present SyntenyPair Explorer, a lightweight, installation-free tool for interactive visualization of synteny between two genomes. The application runs entirely within a standard web browser as a single, self-contained HTML file with no external dependencies and no server-side component. SyntenyPair Explorer accepts standard file formats already produced by common comparative genomics workflows, including FASTA genome assemblies (used to compute optional assembly summary statistics), GFF3/GTF gene annotations, and either BLAST tabular output (outfmt 6) or MCScanX collinearity files. Syntenic relationships are resolved by gene-identifier matching between the relationship file and the gene annotations, so that each relationship corresponds to a discrete gene-to-gene link. Interactive features include continuous zoom and pan, gene search with automatic centering of the partner genome on the syntenic counterpart, synteny block coloring, extensively customizable gene highlights and annotation callouts, session saving and restoration, and publication-quality image export. I demonstrate the tool by visualizing structural differences at the alpha-amylase loci between two strains of the industrially important fungus Aspergillus oryzae. SyntenyPair Explorer lowers the technical barrier to interactive synteny visualization and is freely available under the MIT license at https://github.com/GibbonsLabGenomics/SyntenyPair-Explorer, with a live browser-based version at https://gibbonslabgenomics.github.io/SyntenyPair-Explorer/.

11
Topology-Based Query Framework for Longitudinal Omics Trajectories

Zounemat-Kermani, N.; Richardson, M.; Faiz, A.; Wang, S.; Sun, K.; Vuckovic, D.; van den Berge, M.; Maitland-van der Zee, A. H.; Sayers, I.; Dahlen, S.-E.; Brightling, C. E.; Siddiqui, S.; Chung, K. F.; Nawijn, M. C.; Chadeau-Hyam, M.; Adcock, I. M.

2026-08-21 bioinformatics 10.64898/2026.08.18.745427 medRxiv
Top 0.2%
15.5%
Show abstract

1 Abstract 1.1 Background Many longitudinal omics studies contain only a small number of repeated measurements collected before, during, or after an intervention. Existing approaches, including mixed-effects models and generalized additive models, estimate temporal effects but do not generally provide a discrete representation of trajectory topology that can be queried directly across experimental groups. 1.2 Methods We developed LongOmicsTraj, an open-source R package for topology-based representation and querying of short longitudinal omics trajectories. The framework encodes the direction of change between adjacent visits as up, down, or flat, with the ordered sequence defining an Ordinal Trajectory State (OTS). LongOmicsTraj operates downstream of trajectory estimation and can therefore be applied to empirical summaries or model-derived visit-level estimates, including those from linear mixed-effects models, generalized additive models, and polynomial regression, following a maSigPro-style time-course formulation [1]. OTS labels provide a common representation for topology-based querying, cross-group comparison, and evaluation of higherlevel representations such as trajectory clusters. We evaluated the framework using controlled simulations and bronchial biopsy transcriptomic data from the GLUCOLD corticosteroid intervention study (GEO accession GSE36221), measured at baseline, 6 months, and 30 months. The biological analysis compared continued inhaled corticosteroid (ICS) treatment, ICS withdrawal after 6 months, and placebo. 1.3 Results In simulations, LongOmicsTraj recovered predefined stable, monotonic, transient, rebound, and oscillatory trajectories with high accuracy when longitudinal signal was sufficiently clear, with performance declining under high-noise conditions and depending partly on the upstream estimator. In GLUCOLD, comparator-aware topology queries reduced 20,358 measured transcripts to 168 genes showing a corticosteroid response that was maintained during continued treatment, reversed following withdrawal, and was not reproduced under placebo. The selected genes included established corticosteroid-response genes and were enriched for immune-cell migration, chemotaxis, cell adhesion, and extracellular-matrix organisation. Topology-aware evaluation of FlexMix trajectory clusters additionally revealed substantial within-cluster temporal heterogeneity, with topology purities of approximately 46% to 60%. 1.4 Conclusions LongOmicsTraj provides a compact, directly queryable representation of temporal direction and order in short longitudinal omics studies. It complements existing longitudinal estimation and clustering methods by making trajectory structure explicit, enabling structured cross-group queries and quantification of temporal heterogeneity within trajectory clusters.

12
scDblFinder in Python with GPU support

Hiropedi, A.; Germain, P.-L.

2026-08-20 bioinformatics 10.64898/2026.08.12.744148 medRxiv
Top 0.2%
15.3%
Show abstract

High-throughput single-cell sequencing provides a scalable solution for characterizing cells and profiling gene expression for hundreds to millions of cells. However, this process gives rise to doublets, which can lead to inaccurate conclusions drawn from the data. A number of packages have therefore been developed to help accurately detect them, and in particular scDblFinder has been shown to outperform alternatives in the detection of doublets in single-cell (RNA) sequencing data. Being implemented in R, however, its adoption has been more limited in the Python community. Here, we present scDblFinderPy, a Python-based implementation of the scDblFinder R method, and show that it obtains similar performances. Furthermore, we include in it optional GPU support, thus further speeding up the process.

13
Impacts of batch effects on the performance of machine learning classifiers across multiple studies

Raab, P.; Johnson, W. E.; Piccolo, S. R.

2026-06-30 bioinformatics 10.64898/2026.06.24.734352 medRxiv
Top 0.3%
13.1%
Show abstract

Precision medicine relies on accurate and generalizable predictions for patients across the spectrum of human diversity. Because capturing biological heterogeneity requires large sample sizes, researchers must often aggregate data from several experimental batches or independent studies. This integration allows for greater statistical power and diversity than a single study could provide, while avoiding the costs of generating massive new -omics datasets. Predictive models trained on these aggregated data are theoretically better equipped to detect subtle patterns that generalize to new data. However, this potential is frequently undermined by "batch effects"--systematic technical artifacts that can bias model training to predict experimental batches and shadow meaningful biological conditions. Models trained on data with batch effects can exhibit substantially degraded performance when applied to data from new batches. Statistical adjustment methods can mitigate these artifacts while preserving biological signals. To ensure these adjustments actually facilitate generalization, we emphasize the use of external, independent cohorts for rigorous validation. This chapter examines how batch effects impact predictions and compares various adjustment methods.

14
HydraMPP: A lightweight library for distributed massive parallel processing in Python - threading at scale.

Figueroa, J. L.; White, R. A.

2026-06-08 bioinformatics 10.64898/2026.06.04.730204 medRxiv
Top 0.4%
12.9%
Show abstract

We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC) infrastructures. Massively parallel computing (MPP) has solved this using a divide and conquer approach by splitting workloads across independent nodes (i.e., central processing units (CPU) allowing for higher scaling of data). The main engine for this in python is Ray; however, it has many issues including a large code space, security issues, debugging opacity, and memory management issues. Here, we present HydraMPP, a lightweight, ease of use and utilization, with high auditability, and with SLURM ergonomics.

15
Sequence-Derived Representations versus Pfam-Domain Content for Biosynthetic Gene Cluster Retrieval

Urokov, R.; Khan, A.; Eshboyev, F.; Asadov, D.; Rahman, S.; Kushokova, D.

2026-08-22 bioinformatics 10.64898/2026.08.21.746127 medRxiv
Top 0.4%
12.9%
Show abstract

Retrieving BGCs related to those of a known producer can be regarded as a representation-learning objective. We hypothesize that ESM-2 sequence-derived representations of BGCs can improve retrieval beyond the Pfam-domain content metric. Our toolkit is the following: group-disjoint train, validation, and test assignments, validation-frozen model selection, [fi]ve seeds, and family-level paired inference. Of 6,953 atlas BGCs from 182 deduplicated Streptomyces griseus genome accessions, 5,325 silver-labeled BGCs are split into 98 training, 21 validation, and 21 test reference groups. Of the test reference groups, 16 are eligible for retrieval diagnostics. Pfam Jaccard scored Recall@50 of 0.8788, while Pfam-augmented BGC-SetNet scored 0.8472. The combination of ESM and Pfam-augmented BGC-SetNet scored 0.8769. A weighted Pfam Jaccard obtained a slightly higher score of 0.8789, which has a negligible difference compared to unweighted Pfam accard. Our results do not support the claim that sequence-derived representations can recover alternative biosynthetic pathways on this benchmark. Instead, explicit Pfam remains the major signal for this objective. Our results de[fi]ne the curation and pathway-level validation processes that are necessary for a more robust biological test.

16
Inferring disruption of directed graphs using LIKA reveals altered protein phosphorylation networks in schizophrenia

Zhang, L.; Demarco, A. G.; Ghafari, K.; Devlin, B.; MacDonald, M. L.; Roeder, K.

2026-08-07 bioinformatics 10.64898/2026.08.06.743374 medRxiv
Top 0.4%
12.8%
Show abstract

MotivationKinases regulate a multitude of protein functions, and their dysregulation is pivotal for many human diseases. Direct measurement of kinase activity, however, is often challenging; therefore, inferring activity from the behavior of their substrates is a widely adopted strategy. Nonetheless, traditional methods typically oversimplify the underlying network, ignoring that any particular substrate can be phosphorylated by multiple kinases. ResultsWe present LIKA, a likelihood-based framework for inferring kinase activity from phosphoproteomic data. By modeling the many-to-many structure of kinase-substrate interactions, LIKA achieves high efficiency, even with limited data, while capturing network complexity. Simulation and cell line analyses confirm the robustness and accuracy of LIKA. Importantly, analysis of a phosphoproteomic dataset from schizophrenia and control subjects reveals novel dysregulated kinases. Availability and ImplementationThe implementation code and publicly available data are provided at: https://github.com/lujingz/LIKA.

17
MPGEM: A harmonized and transcriptome-complete resource for large-scale reuse of legacy human microarray data

Gupta, S.; Verma, A. K.; Jana, S.; Ahmad, S.

2026-08-25 bioinformatics 10.64898/2026.08.20.746052 medRxiv
Top 0.4%
12.7%
Show abstract

Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results: We developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource for harmonizing and completing gene-expression profiles across heterogeneous microarray platforms. MPGEM uses a Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework to transform profiles with different gene coverage onto a common quantitative scale. The MPGEM Engine, a multilayer perceptron, predicts expression of unmeasured genes from genes shared across platforms. Applied to Affymetrix GPL570, GPL571, and GPL96, MPGEM uses GPL570 as a 19,320- gene reference space comprising 12,712 predictor and 6,608 target genes. The resulting resource contains 207,135 human gene-expression profiles across 19,320 genes. Evaluation using masked GPL570 profiles yielded mean sample-wise Pearson and Spearman correlations of 0.944 and 0.939, respectively, and mean gene-wise correlations of 0.830 and 0.825. The lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683. MPGEM showed comparable or higher predictive performance than baseline mean imputation and K-nearest-neighbor approaches. Conclusions: MPGEM transforms heterogeneous, partially measured legacy microarray profiles into a harmonized, transcriptome-complete representation, facilitating their reuse for large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning. The framework, trained models, and expression resource are provided as open-source resources.

18
Click-Prep: An Interactive Data Preparation Tool for Click-qPCR

Kubota, A.; Tajima, A.

2026-08-24 bioinformatics 10.64898/2026.08.20.745930 medRxiv
Top 0.4%
12.6%
Show abstract

Click-qPCR is a browser-based application for relative qPCR analysis that requires a tidy-format CSV file containing four columns: sample, group, gene, and Cq. Preparing this input from qPCR instrument output typically requires manual reformatting and calculation of mean Cq values for technical replicates. To simplify this process, we developed Click-Prep (https://kubo-azu.shinyapps.io/Click-Prep/), an interactive web-based application designed specifically to create Click-qPCR input files. Click-Prep imports CSV, TXT, TSV, and XLS/XLSX files and supports skipping of instrument-generated metadata rows, interactive column mapping, and manual assignment of experimental groups. Users can review technical-replicate measurements, exclude selected rows according to predefined quality-control criteria, and calculate mean Cq values for each sample-group-target combination. Missing or nonnumeric Cq values are flagged for review and must be resolved before the mean is calculated. Click-Prep can also combine compatible formatted CSV files, such as datasets obtained from separate qPCR plates. The resulting dataset is exported as a standardized CSV file containing the four fields required by Click-qPCR. By integrating these operations into a guided browser-based workflow, Click-Prep enables users to prepare Click-qPCR input files rapidly and consistently without programming.

19
From Abandoned Scripts to FAIR Community Pipelines: Rescuing Orphan Bioinformatics Workflows with nf-core - Lessons from Light-Sheet Fluorescence Microscopy

Schwitalla, C.; Kuhn Cuellar, L.; Hoertenhuber, M.; Grote, N.; Woller, T.; Lamberti, I.; Pavie, B.; Kuestner, T.; Kyere, F. A.; Curtin, I.; Stein, J. L.; Nahnsen, S.

2026-08-03 bioinformatics 10.64898/2026.07.29.741447 medRxiv
Top 0.4%
12.3%
Show abstract

BackgroundResearch software is essential for modern data analysis but is often developed and maintained by a small number of researchers. When developers leave, software may become orphaned, limiting reuse and risking the loss of valuable domain knowledge and computational methods. While the FAIR Principles for Research Software (FAIR4RS) provide an essential foundation for improving the reuse of research software, compliance with these principles alone does not guarantee practical reusability. Here, we investigate whether orphaned scientific software can be systematically rescued and transformed into sustainable, reusable workflows using established software engineering practices and community standards. FindingsWe re-engineered the abandoned MATLAB-based NuMorph toolkit for large-scale light-sheet microscopy image analysis into nf-core/lsmquant, a Nextflow-based workflow developed according to nf-core community guidelines. The re-engineered workflow preserved the original scientific methods at comparable computational cost while improving the softwares FAIRness, portability, and reproducibility. Integration into the nf-core ecosystem provides a community-driven framework that supports software sustainability through distributed maintenance and shared development practices, while the modular workflow architecture simplified adaptation of nf-core/lsmquant to additional light-sheet microscopy datasets beyond the original application ConclusionOur work demonstrates that orphaned scientific software can be successfully rescued through systematic re-engineering guided by FAIR and software sustainability principles. By transforming a legacy codebase into a community-maintained workflow, we preserve valuable domain-specific methods while improving usability, maintainability, and reproducibility. This approach provides a practical strategy for recovering orphan research software and integrating it into modern, reusable research ecosystems.

20
Semi-automated annotation refinement accelerates cell type identification in brain spatial and single-cell studies

Acri, D. J.; Mustaklem, R.; Horan-Portelance, L.; Dabin, L. C. J.; Park, J. H.; Hartigan, K. A.; Kersey, H. N.; Mesecar, M. E.; Gibbs, J. R.; Cookson, M. R.; Kim, J.

2026-08-04 bioinformatics 10.64898/2026.07.30.741795 medRxiv
Top 0.4%
12.2%
Show abstract

Backgroundsingle-cell and spatial omic techniques have enabled the investigation of cell type specific alterations in biologically complex tissues. In an effort to map cell taxonomies, large atlas-based studies and multi-laboratory consortia have created sets of annotated cell types. However, application of atlas- or database-level knowledge to individual studies is often resource-limited and computational demands scale with the size of both query and reference datasets. ResultsHere, we report a statistical framework for rapid label transfer using summary statistics and user-defined hyperparameters. Semi-Automated Hand Annotation (SAHA)1 allows the user to investigate magnitude, directionality, and statistical significance of matches between unnamed query clusters and reference cell types using either marker-based or marker-free methodologies. By pre-loading the package with summary statistics from the Allen Brain Cell Atlas of the mouse brain, the SAHA R package is capable of rapid cell type comparisons that closely mimic cell typing by integration-based annotation strategies. Furthermore, this flexible package is capable of comparisons across omic modalities, cluster resolutions, and annotations from any study where summary statistics are available. We demonstrate this flexibility by using multiple single-nuclei studies of the mouse cerebellum, mouse cerebral cortex, human cerebral cortex, human peripheral blood mononuclear cells, and one mouse spatial transcriptomic assay. Importantly, this method avoids privacy concerns as it does not require the sharing or deposition of raw data in a web-based tool. ConclusionsAs a result, SAHA offers a non-deterministic annotation reporting structure with automated html reports and summary statistics for transparency in cell typing decisions. Taken together, this scalable framework implemented as a package in R affords increased biological insight into the annotation of single-cell and spatial datasets. SHORT SUMMARYAcri and colleagues present rapid cell type annotation without the need for dataset integration. This paper outlines the utility of the package, SAHA, in annotating neurological datasets.